Papers with extended dataset
Transfer Learning from Transformers to Fake News Challenge Stance Detection (FNC-1) Task (2020.lrec-1)
Copied to clipboard
| Challenge: | In the last two years, significant improvements have occurred in NLP with the development of large language models using contextualized word embeddings based on the Google Transformer architecture. |
| Approach: | They performed experiments on data from the Fake News Challenge stage 1 (FNC-1) they used BERT sentence embeddings as a model feature and BERT, XLNet, and RoBERTa transformers to fine-tune them. |
| Outcome: | The proposed model outperforms the winner's system on class-wise F1 scores and achieves state-of-the-art on the stance detection task. |
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages. |
| Approach: | They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations. |
| Outcome: | The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages. |
Embedding Hallucination for Few-shot Language Fine-tuning (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models can cause severe over-fitting. |
| Approach: | They propose an Embedding Hallucination method which generates auxiliary embedding-label pairs to expand the fine-tuning dataset. |
| Outcome: | The proposed method outperforms current fine-tuning methods in a wide range of language tasks. |
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models (2026.acl-long)
Copied to clipboard
Zirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li, Lang Gao, Zeyu Zhang, Zixiang Xu, Yanbo Wang, Guangxian Ouyang, Zhenhao Chen, Xiuying Chen
| Challenge: | a recent study evaluated large audio-language models against jailbreak attacks . a new benchmark is being developed to evaluate LAM safety against jailbreaking attacks based on temporal and semantic nature of speech . |
| Approach: | They propose a benchmark to evaluate LAM jailbreak vulnerabilities in adversarial audio prompts . they use a dataset of 1,495 adversarials to evaluate their performance . |
| Outcome: | The proposed benchmark evaluates state-of-the-art LAMs against jailbreak attacks . it demonstrates that even small, semantically preserved perturbations can reduce safety . |
Bilingual Zero-Shot Stance Detection (2025.acl-long)
Copied to clipboard
| Challenge: | a study focuses on noun-phrase and claim targets within bilingual ZSSD scenarios . a dataset focusing on claim targets with a low occurrence of shared words is also explored . |
| Approach: | They use a bilingual bilingual ZSSD dataset to investigate the use of zero-shot stance detection. |
| Outcome: | The proposed dataset is the first to examine this difficult setting in bilingual ZSSD . it focuses on noun-phrase and claim targets within in-domain and out-of-domain bilingual scenarios . |